Neural OCR Post-Hoc Correction of Historical Corpora

نویسندگان

چکیده

Abstract Optical character recognition (OCR) is crucial for a deeper access to historical collections. OCR needs account orthographic variations, typefaces, or language evolution (i.e., new letters, word spellings), as the main source of character, word, segmentation transcription errors. For digital corpora prints, errors are further exacerbated due low scan quality and lack standardization. task post-hoc correction, we propose neural approach based on combination recurrent (RNN) deep convolutional network (ConvNet) correct At level flexibly capture errors, decode corrected output novel attention mechanism. Accounting input similarity, loss function that rewards model’s correcting behavior. Evaluation book corpus in German shows our models robust capturing diverse reduce error rate 32.3% by more than 89%.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

OCR and post-correction of historical Finnish texts

This paper presents experiments on Optical character recognition (OCR) as a combination of Ocropy software and data-driven spelling correction that uses Weighted Finite-State Methods. Both model training and testing were done on Finnish corpora of historical newspaper text and the best combination of OCR and post-processing models give 95.21% character recognition accuracy.

متن کامل

Multi-modular domain-tailored OCR post-correction

One of the main obstacles for many Digital Humanities projects is the low data availability. Texts have to be digitized in an expensive and time consuming process whereas Optical Character Recognition (OCR) post-correction is one of the time-critical factors. At the example of OCR post-correction, we show the adaptation of a generic system to solve a specific problem with little data. The syste...

متن کامل

Using SMT for OCR Error Correction of Historical Texts

A trend to digitize historical paper-based archives has emerged in recent years, with the advent of digital optical scanners. A lot of paper-based books, textbooks, magazines, articles, and documents are being transformed into electronic versions that can be manipulated by a computer. For this purpose, Optical Character Recognition (OCR) systems have been developed to transform scanned digital ...

متن کامل

OCR Post-Correction Evaluation of Early Dutch Books Online - Revisited

We present further work on evaluation of the fully automatic post-correction of Early Dutch Books Online, a collection of 10,333 18th century books. In prior work we evaluated the new implementation of Text-Induced Corpus Clean-up (TICCL) on the basis of a single book Gold Standard derived from this collection. In the current paper we revisit the same collection on the basis of a sizeable 1020 ...

متن کامل

Diploma Thesis: Unsupervised Post-Correction of OCR Errors

The trend to digitize (historic) paper-based archives has emerged in the last years. The advantages of digital archives are easy access, searchability and machine readability. These advantages can only be ensured if few or no OCR errors are present. These errors are the result of misrecognized characters during the OCR process. Large archives make it unreasonable to correct errors manually. The...

متن کامل

ذخیره در منابع من

ذخیره در منابع من قبلا به منابع من ذحیره شده

{@ msg_add @}

با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

ژورنال

عنوان ژورنال: Transactions of the Association for Computational Linguistics

سال: 2021

ISSN: ['2307-387X']

DOI: https://doi.org/10.1162/tacl_a_00379